Pseudo-articulatory speech synthesis for recognition using automatic feature extraction from x-ray data
نویسندگان
چکیده
We describe a self-organising pseudo-articulatory speech production model (SPM) trained on an X-ray microbeam database, and present results when using the SPM within a speech recognition framework. Given a time-aligned phonemic string, the system uses an explicit statistical model of co-articulation to generate pseudoarticulator trajectories. From these, parametrised speech vectors are synthesised using a set of artificial neural networks (ANNs). We present an analysis of the articulatory information in the database used, and demonstrate the improvements in articulatory modelling accuracy obtained using our co-articulation system. Finally, we give results when using the SPM to re-score N-best utterance transcription lists as produced by the CUED HTK Hidden Markov Model (HMM) speech recognition system. Relative reductions of 18% in the phoneme error rate and 15% in the word error rate are achieved.
منابع مشابه
Feature Extraction And Sentence Recognition Algorithm In Speech Input System
A feature extraction method for speech waves and an algorithm for sentence recognition are studied. The feature extraction is based on an articulatory model constructed from the statistical analysis of X-ray data. The model holds implicitly the physiological constraints and made possible to estimate the state of the articulatory mechanism. The estimated articulatory parameters provide a set of ...
متن کاملArticulatory Feature Extraction Using CTC to Build Articulatory Classifiers Without Forced Frame Alignments for Speech Recognition
Articulatory features provide robustness to speaker and environment variability by incorporating speech production knowledge. Pseudo articulatory features are a way of extracting articulatory features using articulatory classifiers trained from speech data. One of the major problems faced in building articulatory classifiers is the requirement of speech data aligned in terms of articulatory fea...
متن کاملA Database for Automatic Persian Speech Emotion Recognition: Collection, Processing and Evaluation
Abstract Recent developments in robotics automation have motivated researchers to improve the efficiency of interactive systems by making a natural man-machine interaction. Since speech is the most popular method of communication, recognizing human emotions from speech signal becomes a challenging research topic known as Speech Emotion Recognition (SER). In this study, we propose a Persian em...
متن کاملAutomatic Speech Recognition Based on Electromyographic Biosignals
This paper presents our studies of automatic speech recognition based on electromyographic biosignals captured from the articulatory muscles in the face using surface electrodes. We develop a phone-based speech recognizer and describe how the performance of this recognizer improves by carefully designing and tailoring the extraction of relevant speech feature toward electromyographic signals. O...
متن کاملAn elitist approach to automatic articulatory-acoustic feature classi cation for phonetic characterization of spoken language
A novel framework for automatic articulatory-acoustic feature extraction has been developed for enhancing the accuracy of placeand manner-of-articulation classi cation in spoken language. The ‘‘elitist’’ approach provides a principled means of selecting frames for which multi-layer perceptron, neural-network classi ers are highly con dent. Using this method it is possible to achieve a frame-lev...
متن کامل